Papers with BabyLM challenge

2 papers
LongTail-Swap: benchmarking language models’ abilities on rare words (2025.findings-emnlp)

Copied to clipboard

Challenge: LongTail-Swap is a benchmark that focuses on the tail of the word distribution, i.e., measures the ability of LMs to learn new words with very little exposure, like infants do.
Approach: They introduce LongTail-Swap, a benchmark that measures the ability of language models to learn new words with very little exposure, like infants do.
Outcome: The proposed benchmark measures the ability of language models to learn new words with very little exposure, like infants do.
Is Child-Directed Speech Effective Training Data for Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: High-performing language models are typically trained on hundreds of billions of words, but human learners use language fluently after far less training data.
Approach: They train GPT-2 and RoBERTa models on 29M words of English child-directed speech and a new matched, synthetic dataset.
Outcome: The proposed models show that child language input is not valuable for training language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations